Papers with cognitive tasks
Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation (2025.naacl-long)
Copied to clipboard
Satyapriya Krishna, Kalpesh Krishna, Anhad Mohananey, Steven Schwarcz, Adam Stambler, Shyam Upadhyay, Manaal Faruqui
| Challenge: | Recent advances in Large Language Models (LLMs) have significantly enhanced their capabilities across various cognitive tasks. |
| Approach: | They propose a high-quality evaluation dataset to test LLMs' ability to provide factual responses, assess retrieval capabilities, and evaluate the reasoning required to generate final answers. |
| Outcome: | The proposed framework improves performance in end-to-end RAG scenarios. |
Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs (2026.eacl-long)
Copied to clipboard
Karin de Langis, Jong Inn Park, Bin Hu, Khanh Chi Le, Andreas Schramm, Michael C. Mensink, Andrew Elfenbein, Dongyeop Kang
| Challenge: | Working memory is a critical component of human intelligence and executive functioning . it is correlated with performance on various cognitive tasks, including fluid intelligence . |
| Approach: | They apply working memory tasks to large language models to estimate working memory capacity . they find that LLMs exceed normative human scores, but not executive functioning benchmarks . |
| Outcome: | The proposed models do not show higher performance on executive functioning tasks or problem solving benchmarks. |
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests (2025.findings-emnlp)
Copied to clipboard
Filippo Momentè, Alessandro Suglia, Mario Giulianelli, Ambra Ferrari, Alexander Koller, Oliver Lemon, David Schlangen, Raquel Fernández, Raffaella Bernardi
| Challenge: | MMLU and BBH are three evaluation paradigms for language learning models . interactive games are superior to standard benchmarks in discriminating models based on human cognitive assessments . |
| Approach: | They examine three evaluation paradigms: standard benchmarks, interactive games and cognitive tests . they examine whether interactive games are more effective at discriminating LLMs . |
| Outcome: | The results show that interactive games are superior to standard benchmarks in discriminating models. |
Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for Theory of Mind (ToM) focus on whether agents have correct beliefs about others. |
| Approach: | They propose to evaluate Theory of Mind (ToM) capabilities in Large Language Models (LLMs) they propose to use the theory of mind to determine whether and how to invoke ToM . |
| Outcome: | The proposed frameworks can be used to evaluate the performance of large language models (LLMs) in biological agents. |